Fudan Univers2026-09-15 12:49:09Fudan University Unveils AVTrack Benchmark for Audio-Visual Speaker Tracking, With Paper Accepted to ICML 2026Fudan University’s CVL Lab has introduced AVTrack, a new benchmark aimed at testing audio-visual speaker tracking in complex human-centered scenes. The project paper has been accepted to ICML 2026, and the team has released the project page, paper, code repository, and dataset. AVTrack is built to evaluate whether models can do more than identify who is speaking in a single moment. It measures whether systems can maintain stable speaker identity over time while producing pixel-level masks for active speakers in videos that include motion, occlusion, background shifts, multiple candidates, turn-taking, and mismatches between audio and visible people. The benchmark contains 871 videos averaging 54.0 seconds and provides 3,120 pixel-level instance trajectories with cross-frame identities. In experiments, visual-only methods such as VITA, LBVQ, and CAVIS all scored below 12 on the HOTA test metric. Audio-visual methods improved results, with AVISM at 20.84 and ACVIS at 20.60, while the modular AVTracker baseline reached 29.08. The team also tested Gemini 2.5 Pro in a zero-shot setup and reported a HOTA score of 14.4, suggesting that broad omni-modal understanding does not directly translate into stable pixel-level speaker tracking in difficult scenes.430
Ant Lingbo2026-09-11 10:13:00LingBot-Map lands ECCV 2026 Oral after drawing strong traction on GitHubAnt Lingbo’s streaming 3D reconstruction model LingBot-Map, which was open-sourced in April, has picked up 16.9K GitHub stars and 1.9K forks while repeatedly appearing on GitHub Trending. The project has now received academic recognition as well: its paper, Geometric Context Transformer for Streaming 3D Reconstruction, was selected for an ECCV 2026 Oral presentation, with the work also appearing on the conference’s Award Candidates list seen by Jiqizhixin’s on-site editor at the opening ceremony. The model focuses on a central robotics problem—how a machine can retain spatial memory over long video streams without letting compute and storage costs scale out of control. Using only a standard RGB camera, LingBot-Map performs real-time camera pose estimation and 3D scene reconstruction, and the team says it remains robust across long, multi-room sequences with major viewpoint and environment changes. The article also traces how developers have adapted the model for Apple Silicon Macs, one-click Pinokio deployments, RTX 4060 8GB GPUs, smart glasses workflows, and long-sequence experiments, before detailing the Geometric Context Attention design and the model’s benchmark results on Oxford Spires, ETH3D, 7-Scenes, and NRGBD.920
ECCV2026-09-11 02:51:21ECCV 2026 names best paper, while Fei-Fei Li co-authored work wins test-of-time honorECCV 2026 has announced its annual awards, including best paper, best paper nominees and test-of-time distinctions, after this year’s conference ran from Sept. 8 to 12 in Malmo, Sweden. According to the conference figures cited in the source article, ECCV received 10,473 valid submissions, accepted 2,834 papers for a 27.1% acceptance rate, and selected 163 oral presentations, equal to 1.6%. The best paper award went to Imperial College London’s “Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat,” which introduces HKTex, a mesh-surface texture representation that replaces UV maps with anisotropic heat kernels placed directly on the surface. Two papers were named best paper nominees: Meta Reality Labs Research’s LSRM on scaling context windows for object-centric 3D reconstruction, and Stony Brook University’s Poppy, a training-free framework that uses a single polarization capture at test time to refine surface normal estimation. ECCV also gave three test-of-time awards to “SSD: Single Shot MultiBox Detector,” “Perceptual Losses for Real-time Style Transfer and Super-resolution,” co-authored by Fei-Fei Li, and “Learning Without Forgetting.” The source article also notes that ECVA released information on its PhD Award and Young Researcher Award.410
AI2026-08-12 21:33:05AI-Generated Patterns Shown to Evade Flock Camera Detection in Def Con DemoSecurity researcher Bill Swearingen says he has spent the past year running roughly 31 million tests to build AI-generated patterns that can prevent Flock camera software from recognizing whatever the pattern covers. In a public demo at Def Con, Swearingen worked with YouTube channel Donut Media to wrap a 2009 Toyota Yaris in one of the latest designs and drive it past a Flock camera. The camera still recorded normal footage, but the object-detection layer on top of the video was the part that failed, meaning the system may not classify the vehicle or log its plate. Swearingen said the project, called noRecognition, was built with a reinforcement learning model that keeps adjusting patterns after each failed attempt. He framed the work as a privacy tool for people who want to avoid automated tracking, while also saying the strongest patterns are being kept offline so camera vendors cannot train against them. The project is now running a crowdfunding campaign for merchandise, starting with T-shirts and hoodies and later moving to vehicle wraps.1590
Tsinghua Univ2026-07-24 07:00:17Tsinghua team unveils AutoMIA, an AI system that turns two images into mirror-illusion 3D art for printingResearchers from Tsinghua University and partner institutions have introduced AutoMIA, an AI-based design method for Mirror Illusion Art that can turn any two target images into a single 3D object showing different appearances in direct view and in a mirror. The paper, posted at arXiv and accompanied by open-source code on GitHub, has been accepted to CVPR 2026 as a Highlight Paper and also received the Efficient CVPR award. The system models the task as a dual-view inverse design problem. It represents an object with voxels containing both density and color, then uses differentiable volumetric rendering to optimize geometry and appearance against two target images. The authors said the method addresses four major issues in this kind of design: surface noise, background noise, internal breakage, and imbalance between shape and color optimization. According to the paper, AutoMIA reached 0.989 in smoothness, 0.049 in noise level, and 0.931 in shape similarity, outperforming Shadow Art and Shadow Art Revisited in reconstruction quality. On a single RTX 3090, the average design time was 76 seconds with about 2.6 GB of memory use. The work also includes Blender simulations and 3D-printed physical examples showing that the generated objects can move from digital design into real-world fabrication.280
World Models2026-07-06 06:44:26MemoBench Exposes a Core Weakness in World Models as 10 Video Generators Score Below 0.6 on Object ReappearanceResearchers from Harvard, MIT, IBM, Boston University, Google, Johns Hopkins, CMU, and the Kempner Institute have introduced MemoBench, a new benchmark designed to test whether video generation models can preserve object permanence in dynamically changing environments. The benchmark addresses a major blind spot in current world-model evaluation: most existing tests focus on frame consistency while objects remain visible, but rarely measure whether a model can maintain identity, update state, and restore an object correctly after it leaves the camera view and continues changing off-screen. Built on 360 high-quality ground-truth videos spanning both synthetic and real-world scenes, MemoBench evaluates 10 leading video and world-generation models using automated metrics and VQA-based semantic scoring. The headline result is that no model achieved an object reappearance score above 0.6 out of 1. The findings suggest that visually coherent generation still falls far short of genuine world understanding, especially when memory continuity and state evolution under occlusion are required.740